IEEE Access
● Institute of Electrical and Electronics Engineers (IEEE)
Preprints posted in the last 30 days, ranked by how well they match IEEE Access's content profile, based on 35 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Zhuang, Q.; Mou, C.; Liu, B.; Fu, M. R.; King, G. W.
Show abstract
Breast cancer survivors frequently experience upper-limb impairments, making continuous monitoring essential for effective rehabilitation. We propose REINA (Recognize-Then-Infer Wearable-to-App AI Framework), a two-stage deep-learning approach for remote monitoring of motor function during breast cancer rehabilitation using wearable-device data. Inertial measurement unit (IMU) signals from wearable devices are first used to recognize physical activities via supervised learning, followed by an activity-specific recurrent neural network (RNN) to infer corresponding electromyography (EMG) signals. REINA establishes reliable inference of neuromuscular activity from wearable IMU data, enabling real-time, cost-effective assessment of motor function recovery in real-world settings.
Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.
Show abstract
Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.
Garcia, N. M.
Show abstract
Conventional electrocardiography is highly effective for waveform and rhythm diagnosis, but it is less suited to showing how the internal shape of hundreds or thousands of consecutive heartbeats changes over time. We introduce FOXTAIL, a complementary view that represents each cardiac cycle as an ordered sequence of changes in signal direction. Overlaying these sequences in a fixed visual field makes beat-to-beat organization visible and allows the density, size, stability, and scale persistence of those changes to be measured. We evaluated the representation in recordings containing normal sinus rhythm, paroxysmal atrial fibrillation, severe heart failure, ventricular tachyarrhythmia, and controlled electrode-motion noise. Paired recordings showed that FOXTAIL descriptors can reveal within-person state changes that are not conveyed by a single average beat. The noise and pre-fibrillation analyses also showed that a dense event pattern is not automatically equivalent to physiological complexity, measurement artifact, or impending disease. FOXTAIL is therefore not proposed as a replacement for the diagnostic ECG or as a new classifier, but as an observation and measurement domain for asking a more basic question: how is the electrical organization of the heart changing from one beat to the next, and which of those changes persist across scale?
Ahmed, M.; Otalora, S.; Das Gupta, S.; Kutsuzawa, G.; Akaydin, A.; Le Kernec, J.; Kobayashi, Y.; Mico-Amigo, E.
Show abstract
Prosthesis non-use and abandonment remain common among people with lower-limb amputation, yet current outcome measures capture only limited aspects of how prostheses are used in everyday life. Clinical assessments are typically conducted in controlled settings and rely on self-report or aggregate activity counts, which do not adequately represent functional performance, physiological effort, or lived experience during real-world prosthesis use. Wearable and ambient sensing offer a means of addressing this gap, but existing approaches tend to measure single dimensions in isolation and are rarely validated against laboratory reference standards before free-living deployment. This protocol describes an integrated multimodal framework for assessing real-world lower-limb prosthesis use across three complementary domains: classification of activities of daily living, estimation of energy expenditure, and assessment of emotional state. Approximately 40 adults with unilateral transfemoral or transtibial amputation complete a two-phase protocol. In the laboratory phase, wearable inertial, physiological, and ambient sensing are validated against established reference standards, including video annotation and indirect calorimetry. In the free-living phase, validated models are applied during a single seven-day home monitoring period, unifying all three domains within one deployment. A defined data harmonisation and quality-control procedure aligns heterogeneous sensor streams and preserves traceability between laboratory calibration and free-living measurement, enabling reproducible interpretation of functional behaviour, metabolic cost, and momentary emotional experience in relation to established clinical outcome domains. By integrating multimodal sensing at the level of study design rather than post-hoc analysis, the framework provides a validated, reproducible methodology for characterising prosthesis use beyond the capacity of conventional instruments, offering a transferable approach for real-world monitoring in rehabilitation research
Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.
Show abstract
Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.
Sanz Morere, C. B.; Garrido-Lopez, G.; Hayase, M.; Rueda, J.; An, Q.; Shimoda, S.; Moreno, J. C.; Navarro, E.
Show abstract
Static force plates (FP) are the gold standard for measuring ground reaction forces (GRF) and computing joint moments through inverse dynamics in gait analysis. However, they are restricted to controlled environments, and the number of steps analyzed is limited by the plates embedded in the floor. To address these limitations, portable solutions such as sensorized insoles, socks, or shoes have emerged. Yet, creating wearable systems capable of measuring three-dimensional GRF in real-world conditions remains challenging. Current sensorized shoes often incorporate thick sensors (up to 2 cm), reducing usability and limiting their application in pathological populations or dynamic tasks like running. This study evaluates the usability of ShokacShoes, a novel sensorized shoe integrating three thin, three-dimensional force sensors, and explores its potential as a Wearable Force Plate (WFP). Eight healthy participants performed slow, natural, and fast walking using two insole configurations. Force and temporal metrics were derived from WFP and FP data. Results indicate that WFP enables accurate step segmentation and detects significant effects of speed and insole type on temporal and force metrics, confirming its reliability under different walking conditions. Comparisons with FP revealed differences in force metrics and signal morphology, though temporal parameters remained consistent. These results are likely due to sensor quantity and positioning. Thereby, ShokacShoes represent a valid solution capable of measuring three-dimensional forces within commercial footwear. Future work will focus on validating the applicability of a new version of ShokacShoes against gold-standard FP in a comprehensive validation study involving diverse real-world scenarios and pathological conditions.
Kesenci, Y.; Le Folgoc, L.; Angelini, E.
Show abstract
Deep-learning-based segmentation algorithms have gained considerable accuracy for processing biological images. In particular, the introduction of large foundation models, novel architectures, and semantically varied datasets now allows for deployment of state-of-the-art models for clean image cohorts with limited re-training or, in the best of cases, in an out-of-the-box fashion. Biological imaging, however, is liable to corruptions that can hinder their deployment. While some methods document their robustness to the most common corruptions, a systematic robustness analysis of the state of the art to the expansive gamut of corruptions in biological imaging remains to be done. We perform this benchmarking by simulating 36 corruption types with varying degradation severity on images sampled from 30 different datasets. Our benchmark accounts both for the variety in biological images and the nature of corruptions. Among other things, our study reveals that performance on clean images does not correlate with overall robustness to image corruptions. In fact, we find that a decade-old method, StarDist, is more robust than many of its more recent foundation-model-based counterparts. We also show in a dedicated representation analysis that the performance of segmentation models collapses in the early layers of the encoding phase.
Biasi, N.; Parollo, M.; Vultaggio, D. M.; Zucchelli, G.; Tognetti, A.
Show abstract
Scar-related ventricular tachycardia (VT) is sustained by patient-specific structural and functional remodeling of the ventricular substrate, and the optimal substrate-based ablation strategy remains debated. We present an image-based computational framework for conducting controlled in silico trials of VT ablation strategies. Patient-specific left ventricular electrophysiology models were generated from late gadolinium enhancement cardiac magnetic resonance images by incorporating image-derived scar, border-zone tissue with structural fibrosis, fiber orientation, and physiologically plausible Purkinje-driven sinus activation. A dedicated standalone graphical user interface was developed to perform interactive virtual ablation based on imaging-derived or simulated electrophysiological data. We implemented standardized VT reinducibility testing to compare different lesion sets in terms of residual VT inducibility and ablation burden. As a proof of concept, the framework was applied to 20 patients with ischemic or non-ischemic cardiomyopathy undergoing VT ablation. Sustained VT was inducible in 17 patients, yielding 127 sustained VT episodes and 88 unique reentrant circuits at baseline. Four substrate-based ablation strategies were compared: scar homogenization, primary deceleration-zone ablation, primary plus secondary deceleration-zone ablation, and CMR-guided scar dechanneling. All strategies significantly reduced VT inducibility compared with baseline. Scar homogenization achieved the largest reduction in residual unique sustained VTs but required the largest ablated myocardial volume. Conversely, CMR- guided scar dechanneling showed the most favorable efficiency profile by reducing VT inducibility while limiting ablated viable myocardium. The proposed framework enables quantitative comparison of ablation efficacy, ablation burden, and mechanisms of ablation success or failure in image-guided VT therapy planning.
Wegner, P.; Ophey, A.; Roettgen, S.; Kufer, K.; Doppler, C. E.; Seger, A.; Fink, G. R.; Kalbe, E.; Kotra, K.; Grobe-Einsler, M.; Feldmann, K.; Sommerauer, M.; Faber, J.
Show abstract
Objective and scalable approaches for detecting subtle motor impairment in isolated REM sleep behavior disorder (iRBD), a prodromal stage of Parkinson's disease, remain limited. We investigated whether markerless motion capture from single RGB-camera videos can identify gait abnormalities in people living with iRBD and provide interpretable digital biomarkers. We retrospectively analyzed 93 standardized walking videos from three clinical sites. Human pose estimation extracted 12 body markers and 14 kinematic time series. Thirty-five machine learning approaches classified healthy controls (HC) and people with iRBD. The Movement Disorder Society Unified Parkinson's Disease Rating Scale Part 3 (MDS-UPDRS III) served as the clinical baseline. The best-performing model (tsfresh+XGBoost) achieved an AUROC of 0.739, significantly outperforming the MDS-UPDRS III sum score when trained on data from all three sites. Harmonized multi-site training improved performance. SHAP identified hip-related temporal features as key contributors, which differed between groups and showed stronger associations with regional dopaminergic deficits than clinical scores. Single-camera gait analysis may provide scalable digital biomarkers for low-cost screening and monitoring of prodromal PD.
Dack, E.; Dai, C.; Hoppe, H.; Krueselmann, P.; Meiler, S.; Jutidamrongphan, W.; Wang, L.; Tang, K.
Show abstract
AI-assisted diagnostic tools typically act as a "second opinion," providing radiologists with a discrete prediction or probability score that can be consulted alongside clinical context. This treats AI as an independent advisor rather than a collaborative partner, leaving its reasoning largely opaque. We explore a complementary approach grounded in human-AI collaboration through visual interpretability. Specifically, we investigate (1) radiologist performance when diagnosing chest X-rays from images alone, and (2) whether deep learning-generated heatmaps can support radiologists during this diagnostic process, rather than merely validating a final answer. We developed an interactive application that enables readers to engage directly with model-generated heatmaps as they form their diagnoses, and conducted a user study to evaluate how this influences diagnostic behaviour and accuracy. Our findings offer new insights into integrating interpretable, spatially grounded AI feedback into radiologist workflows. Code, datasets, and the application can be found at https://github.com/eedack01/heatmap_assisted_diagnosis.
Bai, X.; Kishimoto, K.; Sugiyama, O.; TAMURA, H.
Show abstract
This study aims to improve the detection performance of age-related macular degeneration (AMD) in low-quality retinal images. BackgroundAMD is a leading cause of vision loss among older adults globally, and accurate detection is crucial for clinical management. However, low-quality optical coherence tomography (OCT) images significantly compromise diagnostic accuracy. ObjectiveTo enhance AMD detection in low-quality images using noise-augmented data augmentation and an improved YOLO deep learning model. MethodsPublic datasets from UCSD and Duke University were utilized; the training dataset comprised 24,980 OCT images (high-quality and noise-augmented low-quality), while the testing dataset included 1,000 images (584 AMD, 416 normal). The model is based on the YOLOv8n framework, integrated with Squeeze-and-Excitation blocks (SEblock) and Adaptive Sparse Self-Attention (ASSA), with an additional 160x160 detection layer for detecting small lesions. Evaluation metrics included accuracy, sensitivity, specificity, and F2-score. ResultsThe proposed model achieved an accuracy of 99.02%, sensitivity of 98.17%, specificity of 100%, and an F2-score of 98.50% on the Duke dataset. Detection rates were significantly improved compared to traditional methods, particularly in low-quality images, with a detection rate of 89.60%, markedly superior to original YOLOv8n (55.10%) and classical models like ResNet50. ConclusionThe enhanced model, employing noise-augmented training data and improved attention mechanisms, demonstrates excellent AMD detection capabilities in low-quality OCT images, showing broad potential for clinical applications.
Liu, D.; Dutta, A.; Nadig, S.
Show abstract
The features of the PPG (photoplethysmography) morphology are known to reflect age-related cardiac and vascular changes. In most contemporary wearables, PPG signals are acquired from distal sites such as the wrist and finger. The superficial temporal artery (STA), accessible at the temple region, is reached via a shorter arterial path from the aortic root than the radial circulation, and may therefore carry hemodynamic and aging information with less distance-dependent attenuation. We hypothesized that the morphology of the PPG at temple region (STA) would show stronger and more numerous age correlates than the PPG at the wrist. To test this, we extracted a common set of 89 pulse-morphology features, spanning raw-waveform timing/amplitude/area measures, ratios among them, derivative-based ratios, and spectral harmonic-ratio features. We compared an in-house temple-worn device which has PPG as one of the sensors, with a publicly available Microsoft Aurora-BP wrist-worn PPG dataset, and tested each feature's association with age. We identified 14 robust age correlates at the temple region, compared to 3 at the wrist. The temple's correlates spanned multiple morphological categories and showed a larger age-association than at the wrist. These results support the hypothesis that the temple region may be a more robust PPG measurement site than the wrist to extract age-related cardiovascular information, which motivates further investigation of temple-based cardiovascular sensing.
Kumar, B. R.; Ramsundar, B.; Subramanian, S.
Show abstract
Neural temporal point processes (NTPPs) are powerful tools for modeling sequences of timestamped events with statistical temporal structure. Density-based NTPPs, in particular, are an interesting opportunity to merge the universal function approximation capability of neural networks with a defined statistical model in a way that has many potential applications. We demonstrate one such application to heartbeat dynamics, a physiologic point process. We specifically apply a lognormal mixture NTPP to compute instantaneous estimates of the mean and standard deviation of beat-to-beat intervals. We compare our results to the state of art (Barbieri et al.) point process model for heartbeat dynamics, which uses a more physiologically rigorous inverse Gaussian model. We find that the NTPP model maintains reasonable accuracy while improving upon robustness to noise.
Robbins, C.; Son, H.; Tan, C. K.; Wang, C.; van Kanten, R.; Sartori, M.; Durandau, G.; Kumar, V.; Caggiano, V.; Song, S.
Show abstract
Physical human-device interaction is central to many emerging technologies in neurorehabilitation and assistive robotics, but simulation-based research in this area remains fragmented across musculoskeletal models, assistive-device representations, task definitions, and controller-development workflows. This fragmentation limits the accessibility, reproducibility, and extensibility of studies on prostheses, exoskeletons, wearable rehabilitation devices, and related human-device systems. Here we introduce MyoAssist 1.0, an open-source framework for neuromechanical simulation of physical human-device interaction built within the MyoSuite ecosystem. MyoAssist organizes each simulation environment as a composed human-device-task system that combines compatible musculoskeletal, assistive-device, and task-scenario components through a shared composition pipeline. The current release includes 15 assistive-device models spanning gait assistance, upper-body support, manipulation, and seated mobility and supports compatible musculoskeletal models ranging from reduced lower-limb models to a 416-muscle full-body model. These human-device systems can be simulated within the broad task scenarios provided by MyoSuite, while MyoAssist adds locomotion-specific task scenarios with configurable terrain and target-velocity conditions for gait-assistive studies. MyoAssist also provides two complementary controller-development frameworks: a reinforcement-learning framework for training adaptive policies and a controller-optimization framework for tuning structured, interpretable human and device controllers. Both frameworks operate on the same simulation environments and provide standardized evaluation outputs for inspecting, comparing, reusing, and extending learned and structured control strategies. By integrating modular human models, assistive-device models, task scenarios, and training workflows under a shared open-source interface, MyoAssist aims to lower the barrier to reproducible simulation-based research and to support collaborative development of assistive technologies for neurorehabilitation and physical human-device interaction.
Devatha, D.; Xiao, J.
Show abstract
Dementia affects more than 55 million people worldwide, and its progressive decline is difficult to track using infrequent in-person assessments, which can miss subtle changes between visits and add to clinician burden. One widely used measure, the Mini-Mental State Examination (MMSE), is administered intermittently and remains subject to inconsistent scoring judgment, limiting early detection. Prior speech-based machine learning approaches have largely focused on cross-sectional classification rather than longitudinal cognitive forecasting. We introduce a longitudinal, patient-level framework that combines transformer-derived semantic speech representations with longitudinal speech-change features and clinical history to forecast a patient's future MMSE score from their history of prior visits. To our knowledge, this is the first framework to unite transformer-derived speech encoding with longitudinal forecasting of cognitive severity, rather than single-visit classification alone. We evaluate this framework on longitudinal transcripts from the DementiaBank Pitt Corpus using a LightGBM gradient-boosted regression model, validated with patient-grouped cross-validation to prevent identity leakage between training and evaluation folds. The model forecasts future MMSE scores with high accuracy and stability across folds (R^2 = 0.840 +- 0.015, RMSE = 2.75, MAE = 2.02, Pearson r = 0.917), with transformer-derived speech representations contributing meaningful predictive signal alongside clinical history. Ablation analysis demonstrated that transformer-derived speech representations provided complementary predictive information beyond clinical variables. These results establish speech as a viable longitudinal digital biomarker of cognitive decline, offering a low-burden complement to intermittent clinical assessment that could enable earlier detection and more frequent monitoring, supporting better-timed care decisions for patients with dementia.
Palangattu, A.; Sah, A. K.; Raman, S.; Pushpavanam, K. S.
Show abstract
In materials science, the integrity of scanning electron microscopy (SEM) images is paramount for quality control and validation of research outcomes. However, the introduction of sophisticated generative artificial intelligence, particularly Generative Adversarial Networks (GANs), has introduced a novel vulnerability: the potential for highly realistic, artificially synthesized SEM images to be used fraudulently in scientific literature. To address this challenge, we present a deep learning-based framework capable of distinguishing between authentic SEM images and those synthesized by Generative Adversarial Networks (GANs). Using FastGAN and StyleGAN2-ADA, two state-of-the-art GAN models, we generated synthetic SEM datasets to complement real imaging data. We fine-tuned a pre-trained Contrastive Language-Image Pre-training (CLIP) Vision Transformer (ViT-L-14) for binary classification. By unfreezing the final transformer blocks and appending a custom classification head, the model effectively captures the subtle, high-level artifacts inherent in GAN-generated upsampling. This work highlights the potential of deep learning to safeguard scientific imaging workflows and provides an important step toward detecting and mitigating image forgeries in materials science publications.
Mardaljevic, J.; de Vries, S. W.; van Duijnhoven, J.
Show abstract
The measurement of light received at the cornea of the eye is a paramount consideration for the understanding of the relation between environmental illumination and the non-image-forming effects of light. The field of view (FOV) at the cornea is less than a full hemisphere, because it is partially occluded by human facial morphology. The International Commission on Illumination (CIE) has defined a standard model of human FOV. A suitably designed physical occluder attached to the sensor (of a light meter) has been proposed as a means of incorporating the effect of human FOV when taking measurements. Similarly, when using simulation to predict light received at the cornea, a geometrical description of the occluder at the eye point(s) can be added to the 3D model of the scene. The first occluder model proposed to represent CIE human FOV was enumerated in terms of: the CIE definition; the radius of the occluder; and, the radius of the light sensor disc. We present a simpler model based only on the CIE definition and the occluder radius. Both models were tested using a virtual goniophotometer. Various sensor response functions describing the spatial sensitivity across the sensor disc, including several we characterized through laboratory measurements, were included in the test. For all functions considered, the performance of the simpler occluder model was equivalent to or better than the model first proposed.
Sunil, G.; Kumar, B. R.; Ramsundar, B.; Subramanian, S.
Show abstract
Scaling laws help determine the optimal data size for training large models but are established in domains where the target is deterministic. Physiological signals are different: heartbeat sequences are stochastic, so part of the error is irreducible even with large amounts of data. Metrics such as MAE do not account for non-deterministic behavior, and therefore assessing scaling requires evaluating distributional calibration (measuring how well predicted probability densities capture true conditional characteristics). We formulate a scaling law metric(n) = E + A n- and evaluate it with five metrics: accuracy (MAE, RMSE), distributional calibration (KS distance, goodness-of-fit), and training objective (negative log loss) using a neural temporal point process trained on a cohort of four-ECG datasets. The law fits all five metrics. While point accuracy is near saturation at n = 183, KS distance and goodness-of-fit improve by 6% and 12% respectively when extrapolated to 10,000 subjects, showing that scaling decisions in stochastic domains must be guided by distributional calibration rather than point accuracy.
Ye, Z.; He, F.; Zhao, T.; Xia, W.
Show abstract
Ultrathin endoscopy is highly attractive for real-time tissue imaging in narrow and hard-to-reach regions of the body. A single multimode fibre (MMF) is an attractive probe because of its small diameter, flexibility, and diffraction-limited spatial resolution enabled by the large number of transverse modes guided within a single core. Because the distal fibre tip is inaccessible during endoscopy, reflection-mode imaging, in which the same fibre delivers illumination and collects backscattered light, is more practical than transmission-mode imaging. However, image recovery from the resulting speckle pattern is challenging because light undergoes double-pass propagation through the MMF, with mode coupling and dispersion; the backscattered signal is weak, and the camera records intensity only, without phase information. Here, we propose a single-shot reflection-mode MMF imaging framework that combines a reflected real-valued intensity transmission matrix (reflected-RVITM) with an image restoration network. The reflected-RVITM is calibrated using intensity-only measurements, without interferometry or phase retrieval, and provides a physics-guided initial reconstruction from a single backscattered speckle frame. A restoration network then refines this initial reconstruction instead of inverting the raw speckle. Four restoration backbones are evaluated: HPM-Attention-UNet, GAM, MambaIRv2, and CICPNet. On matched datasets, hybrid models outperformed corresponding networks trained to map raw speckle directly to images. For example, HPM-Attention-UNet on MNIST improved mean PCC from 0.572 to 0.944 (+65.1%). Under domain shift, with training only on Fashion-MNIST and tested on unseen CIFAR scenes, hybrid models achieved mean PCC of 0.61-0.65, compared with 0.36-0.50 for direct learning. This framework is further demonstrated using physical objects at the distal fibre tip. These results demonstrate that a reflected-RVITM physics prior combined with a restoration network enables single-shot image recovery after intensity-only calibration, offering a phase-retrieval-free and generalisable route towards minimally invasive reflection-mode MMF endoscopy.
Nayak, K. S.; Nirgude, A. S.; Das, R.
Show abstract
Background Stroke remains one of the leading causes of mortality and long-term disability worldwide, with low- and middle-income countries bearing a disproportionate share of the global disease burden. In India, delays in risk identification, fragmented referral pathways, and limited continuity of preventive care present significant challenges, particularly in rural communities. As a frontline health worker Accredited Social Health Activists (ASHAs) are strategically positioned to support community-based stroke prevention; however, existing workflows are frequently constrained by multi-tasking, predominantly paper-based documentation and fragmented digital systems. Advances in mobile health, artificial intelligence along with digital health ecosystem provided by Ayushman Bharat Digital Mission (ABDM) provide an opportunity to strengthen community healthcare through integrated digital platforms. Objective This protocol describes the design, system architecture, and prospective evaluation framework of ASHA Assist India, an integrated AI-assisted mobile health platform intended to support community-based stroke prevention by connecting citizens, ASHA workers, Primary Health Centres (PHCs), and higher levels of healthcare facilities within a unified digital ecosystem. Methods ASHA Assist India has been designed as a modular, cloud-based digital health platform supporting standardized data collection, longitudinal health monitoring, referral management, and AI-assisted clinical decision support. The proposed system comprises four user-facing applications corresponding to citizens, ASHA workers, PHCs, and referral hospitals, integrated through a centralized backend providing authentication, secure data management, interoperability, analytics, and notification services. The AI framework includes three planned analytical modules: (i) population-level stroke risk stratification, (ii) longitudinal stroke risk prediction, and (iii) acute stroke symptom recognition. A prospective implementation study is planned to evaluate platform usability, feasibility, workflow integration, implementation outcomes, and operational performance within routine community healthcare settings. Future validation of the AI modules will be conducted using prospectively collected longitudinal datasets. Expected Impact The proposed platform aims to strengthen community-based stroke prevention by improving digital workflow integration, facilitating coordinated referral pathways, and supporting longitudinal monitoring through the existing healthcare providers at health and wellness centres like ASHA, Community Health Officers (CHOs), ANM, etc. Beyond stroke prevention, the modular architecture is intended to provide a scalable framework for future digital health programmes addressing multiple non-communicable diseases within primary healthcare systems. Publication of this protocol establishes a transparent implementation and evaluation framework that may guide future research, digital health innovation, and implementation science in resource-constrained settings.